Infrastructure
Infrastructure decisions in health are shaped less by technical preference than by three constraints that do not apply elsewhere: data residency law, availability requirements that are clinical safety requirements, and operational capacity that is often the binding limit.
The most sophisticated architecture is worthless if there is nobody to operate it at 3 a.m.
Hosting models
| Model | Suits | Costs |
|---|---|---|
| On-premises (facility) | Poor connectivity; data must not leave the building | Every facility needs power, hardware, backup and someone to maintain them |
| On-premises (national data centre) | Residency requirements; existing government capacity | Capital-intensive; scaling is procurement, not a configuration change |
| Private cloud | Government cloud programmes; residency with elasticity | Depends on the provider's maturity, which varies enormously |
| Public cloud | Elastic demand, managed services, mature operations available | Residency and sovereignty; egress costs; skills; currency and budget volatility |
| Hybrid | The common real answer — central services in cloud, facility systems local | Two operating models to sustain |
| Edge | Facilities that must function while disconnected | Sync, conflict resolution, device management at scale |
Data residency is usually the first filter, and it is a legal question, not a technical one. Establish what the law and the health ministry's policy actually require — often "in-country" for identifiable data, with more latitude for aggregates — before evaluating providers. See GDPR where it applies, and national data protection law everywhere.
A frequent mistake: choosing a global public cloud region in a neighbouring country because it is nearest, without checking whether that is lawful. It usually is not, for identifiable clinical data.
Availability as a clinical parameter
Availability targets should be set by clinicians against consequences, not chosen from a list.
| System | Realistic target | Because |
|---|---|---|
| Emergency department EMR | Very high; minutes of downtime matter | Care stops |
| Facility EMR | High during operating hours | Care degrades to paper |
| Interoperability layer | High — everything routes through it | Blocks all exchange |
| Client registry | High — on the registration critical path | Blocks registration |
| Shared health record | Medium; degraded operation is tolerable | Clinicians work from local records |
| HMIS / reporting | Lower; hours of downtime are acceptable | Reports are late |
| Analytics / warehouse | Lowest | Nobody is harmed |
Two consequences for design. First, tiering: not everything needs the same investment, and pretending otherwise means nothing gets it. Second, degraded mode — what the facility does when the central service is unreachable — is an architecture requirement for every high-tier system. Local caching of registries and store-and-forward queues at the edge are the standard answers. See offline-first.
Containers and orchestration
Docker has become the default packaging format for health platforms — OpenMRS, DHIS2, HAPI FHIR, OpenHIM and most others publish images. That alone is a large improvement over hand-built servers.
Kubernetes is a separate decision, and a heavier one:
Adopt it when you run many services with differing scaling needs, have multiple teams deploying independently, need self-healing and rolling upgrades, and — critically — have people who can debug it.
Do not adopt it when you run three services on two servers with one system administrator. Docker Compose on well-managed virtual machines is a legitimate production architecture for a great many national deployments, and it is recoverable at 3 a.m. by someone who did not build it.
The failure mode is a ministry that adopts Kubernetes because it is best practice, and then cannot diagnose a failing pod during an outage. Operational capacity is a design input, not an aspiration.
Managed Kubernetes removes the control-plane burden but not the application operations burden, which is where most of the difficulty actually is.
Compute and storage patterns
| Need | Options | Health notes |
|---|---|---|
| Application runtime | VMs, containers, serverless | Serverless suits event handlers and bulk processing; poor fit for stateful clinical platforms |
| Relational storage | PostgreSQL (usual), MySQL | Most health platforms assume one specifically; check before choosing |
| Object storage | S3-compatible | DICOM images, bulk exports, backups |
| Search | Elasticsearch / OpenSearch | Patient search, terminology search, log analytics |
| Caching | Redis / Valkey | Terminology expansions, session state |
| Messaging | Kafka, RabbitMQ, NATS | Event-driven exchange; see patterns |
Storage growth is dominated by imaging. A radiology programme changes the infrastructure cost profile by an order of magnitude, and retention periods are measured in decades — tiered storage and a lifecycle policy are requirements, not optimisations.
Networking
- Segmentation. Clinical systems, medical devices, administrative networks and guest wifi must be separate. Medical devices are frequently unpatchable; isolation is the only available control.
- Facility connectivity. Design for what exists — intermittent, low bandwidth, high latency — rather than for what is promised.
- Private connectivity between the national data centre and cloud, where hybrid.
- Egress control. Preventing data leaving is as important as preventing entry, and is more often overlooked.
- DNS and certificates. Both are recurring outage causes. Automate renewal and monitor expiry; see security architecture.
Cloud providers
Documented here for completeness. This knowledge base is not vendor-specific, and no provider is recommended over another.
- AWS — HealthLake, HealthImaging, and general services
- Azure — Azure Health Data Services, Microsoft Cloud for Healthcare
- GCP — Cloud Healthcare API
- Oracle Cloud, DigitalOcean, Hetzner, OVH and regional providers — often relevant where residency, cost or existing government agreements dominate
The managed FHIR/DICOM services are genuinely useful and create genuine lock-in. Evaluate the exit path — can you export everything in a standard format, and what would running the equivalent yourself cost? — as part of the selection, not afterwards.
In this section
- Docker — containers and deployment
- Observability — logs, metrics, traces, SLOs
- DevSecOps — CI/CD, infrastructure as code, supply chain
- AWS, Azure, GCP
A sizing checklist
- Data residency requirements established, in writing, from legal counsel
- Availability tier assigned per system, agreed with clinical leadership
- Degraded mode designed for every high-tier system
- RTO and RPO per system, set by consequence
- Backup strategy including one immutable copy, with tested restores
- Growth projection including imaging and audit logs
- Operational team identified and staffed for the chosen complexity
- Monitoring and alerting in place before go-live, not after
- Disaster recovery site or region, and a rehearsed failover
- Ten-year cost model including staff, not only hosting
References
- WHO Digital Health Platform Handbook — https://www.who.int/publications/i/item/9789240013728
- CIS Benchmarks — https://www.cisecurity.org/cis-benchmarks
- Kubernetes — https://kubernetes.io/
- NIST SP 800-207 Zero Trust Architecture — https://csrc.nist.gov/pubs/sp/800/207/final